Papers with speech quality

12 papers
Automated Cross-language Intelligibility Analysis of Parkinson’s Disease Patients Using Speech Recognition Technologies (P19-2)

Copied to clipboard

Challenge: PD is the second most common neurodegenerative disorder after Alzheimers disease . speech impairments are one of the earliest manifestations in PD patients .
Approach: They propose to analyze the speech signals of PD patients and healthy control subjects in three different languages: German, Spanish, and Czech.
Outcome: The proposed model can discriminate between PD patients and HC subjects even when the language used for train and test is different.
A Fast and High-quality Text-to-Speech Method with Compressed Auxiliary Corpus and Limited Target Speaker Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to generate high-quality speech with limited target speaker corpus require extensive training data.
Approach: They propose an auxiliary corpus compression algorithm that reduces the training cost while the naturalness of synthesized speech is not significantly degraded.
Outcome: The proposed method significantly reduces training costs while maintaining the naturalness of synthesized speech.
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer (2024.acl-long)

Copied to clipboard

Challenge: Experimental results indicate that GPST significantly outperforms the existing speech language models in terms of word error rate, speech quality, and speaker similarity.
Approach: They propose a hierarchical transformer that quantizes audio waveforms into two distinct types of discrete speech representations and integrates them within a transformer architecture.
Outcome: The proposed model outperforms existing speech language models in word error rate, speech quality, and speaker similarity.
MOSPC: MOS Prediction Based on Pairwise Comparison (2023.acl-short)

Copied to clipboard

Challenge: et al., 2016a) show that MOS prediction model can improve ranking accuracy of speech quality.
Approach: They propose a general framework for MOS prediction based on pair comparison . they use C-Mixup algorithm to enhance generalization performance of MOSPC .
Outcome: The proposed model outperforms baselines on most correlation coefficient metrics . it also surpasses the strong baseline in ranking accuracy on each fine-grained segment.
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation (2026.eacl-long)

Copied to clipboard

Challenge: Current speech evaluation systems rely on specialized systems for individual audio characteristics and poor correlation between automatic methods and human preferences.
Approach: They propose a unified evaluation framework for Large Audio Models as a Judge, AudioJudge . they propose specialized judges that can be prompted to perform audio characteristic detection tasks .
Outcome: The proposed method improves performance across audio characteristic detection and human preference simulation tasks.
Scaling Rich Style-Prompted Text-to-Speech Datasets (2025.emnlp-main)

Copied to clipboard

Challenge: Existing datasets that only cover basic tags are limited in their scale or coverage of style tags.
Approach: They propose a large-scale dataset that annotates speech utterances with rich style captions.
Outcome: The proposed dataset scales speech utterances with rich style captions for the first time.
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion (2025.naacl-long)

Copied to clipboard

Challenge: Recent advances in text-to-speech (TTS) models have led to improvements in speaker prosody and voices modeling.
Approach: They propose an efficient zero-shot TTS model that leverages distilled time-varying style diffusion to capture diverse speaker identities and prosodies.
Outcome: The proposed model surpasses state-of-the-art models in both naturalness and similarity while reducing inference speed by 90%.
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech (2024.acl-long)

Copied to clipboard

Challenge: Existing zero-shot text-to-speech systems require a few seconds of unseen speaker voice prompts to generate high-quality voices.
Approach: They propose a zero-shot text-to-speech system based on mobile devices . they use a discrete speech codec to integrate hierarchical information from the codec .
Outcome: The proposed system achieves RTF of 0.09 on a single A100 GPU and has been successfully deployed on mobile devices.
Self-EmoQ: Plutchik-Guided Value-based Planning to Drive Streaming Emotional TTS (2026.findings-acl)

Copied to clipboard

Challenge: Existing systems lack a self-emotion determination mechanism to drive the streaming text-to-speech (TTS) synthesis.
Approach: They propose an emotion-planning framework that determines the emotion prior to the textual generation, grounding the downstream emotional TTS in a streaming manner.
Outcome: The proposed framework outperforms baselines on DailyDialog, EmoryNLP, IMEOCAP, and MELD on emotional alignment, contextual coherence, and expressive fluency.
VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show VocalNet outperforms existing open-source speech LLMs despite limited training data.
Approach: They propose a scalable and model-agnostic training framework and a novel multi-token prediction paradigm for speech generation.
Outcome: The proposed model outperforms open-source speech LLMs while outperforming existing open-sourced models.
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM (2025.findings-acl)

Copied to clipboard

Challenge: Existing speech-enabled LLMs degrade conversational quality by modifying the LLM, compromising its linguistic capabilities.
Approach: They propose a lightweight 30M-parameter, LLM-agnostic, autoregressive streaming TTS system that generates high-quality speech with low latency.
Outcome: The proposed system achieves a significantly lower word error rate compared to speech-enabled LLMs while operating at comparable latency.
DM-Codec: Distilling Multimodal Representations for Speech Tokenization (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing speech tokenization models lack contextual representations for speech synthesis . absence of contextual representation results in elevated WER and WIL scores .
Approach: They propose a language model-guided distillation method that incorporates contextual information into a comprehensive speech tokenizer.
Outcome: The proposed method outperforms state-of-the-art tokenization models in reducing WER and WIL scores.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations